Papers by Yash Mahesh Bangera
Language Variety Identification with True Labels (2024.lrec-main)
Copied to clipboard
Marcos Zampieri, Kai North, Tommi Jauhiainen, Mariano Felice, Neha Kumari, Nishant Nair, Yash Mahesh Bangera
| Challenge: | Language identification datasets are compiled with the assumption that the gold label of each instance is determined by where texts are retrieved from. |
| Approach: | They present a human-annotated multilingual dataset for language variety identification . they use a model to train multiple models to discriminate between different languages . |
| Outcome: | The proposed dataset provides a reliable benchmark toward robust and fairer language variety identification systems. |